Gemma 4

Our most intelligent open models, built from Gemini 3 research and technology to maximize intelligence-per-parameter

Model sizes


Performance

Industry-leading efficiency

A scatter plot titled 'Model Performance VS Size' that maps Elo Score (y-axis, from 1350 to 1460) against Total Model Size in Billion Parameters (x-axis, on a logarithmic scale from 10 to over 1000). A blue-shaded region in the top-left corner highlights two relatively small but high-performing models: 'gemma-4-31B-thinking' (at approximately 31B parameters with an Elo score of 1451) and 'gemma-4-26B-A4B-thinking' (at approximately 26B parameters with an Elo score of 1442). Other models, plotted as grey points, include 'qwen3.5-27b' (Elo 1404), 'gpt-oss-120b' (Elo 1354), 'qwen3.5-122b-a10b' (Elo 1416), 'qwen3.5-397b-a17b' (Elo 1450), 'mistral-large-3' (Elo 1416), 'deepseek-v3.2-exp-thinking' (Elo 1425), 'glm-5' (Elo 1456), and 'kimi-k2.5-thinking' (Elo 1455), showing that the highlighted Gemma models achieve Elo scores comparable to those of much larger models.A scatter plot titled 'Model Performance VS Size' that maps Elo Score (y-axis, from 1350 to 1460) against Total Model Size in Billion Parameters (x-axis, on a logarithmic scale from 10 to over 1000). A blue-shaded region in the top-left corner highlights two relatively small but high-performing models: 'gemma-4-31B-thinking' (at approximately 31B parameters with an Elo score of 1451) and 'gemma-4-26B-A4B-thinking' (at approximately 26B parameters with an Elo score of 1442). Other models, plotted as grey points, include 'qwen3.5-27b' (Elo 1404), 'gpt-oss-120b' (Elo 1354), 'qwen3.5-122b-a10b' (Elo 1416), 'qwen3.5-397b-a17b' (Elo 1450), 'mistral-large-3' (Elo 1416), 'deepseek-v3.2-exp-thinking' (Elo 1425), 'glm-5' (Elo 1456), and 'kimi-k2.5-thinking' (Elo 1455), showing that the highlighted Gemma models achieve Elo scores comparable to those of much larger models.
BenchmarkGemma 4
31B IT ThinkingGemma 4
26B A4B IT ThinkingGemma 4
E4B IT ThinkingGemma 4
E2B IT ThinkingGemma 3
27B IT
Arena AI (text) As of 4/2/26145214411365
MMMLU Multilingual Q&ANo tools85.2%82.6%69.4%60.0%67.6%
MMMU Pro Multimodal reasoning76.9%73.8%52.6%44.2%49.7%
AIME 2026 Mathematics No tools89.2%88.3%42.5%37.5%20.8%
LiveCodeBench v6 Competitive coding problems80.0%77.1%52.0%44.0%29.1%
GPQA Diamond Scientific knowledgeNo tools84.3%82.3%58.6%43.4%42.4%
τ2-bench Agentic tool useRetail86.4%85.5%57.5%29.4%6.6%

These models were evaluated against a large collection of datasets and metrics to cover different aspects of text generation. See additional benchmarks in model card.


E2B and E4B

A new level of intelligence for mobile and IoT devices

Audio and vision support for real-time edge processing. They can run completely offline with near-zero latency on edge devices like phones, Raspberry Pi, and Jetson Nano.


12B, 26B, 31B

Frontier intelligence, efficient performance

Advanced reasoning for IDEs, coding assistants, and agentic workflows. These models are optimized for consumer GPUs — giving students, researchers, and developers the ability to turn workstations into local-first AI servers.


Safety

Gemma 4 models undergo the same rigorous infrastructure security protocols as our proprietary models. By choosing Gemma 4, enterprises and sovereign organizations gain a trusted, transparent foundation that delivers state-of-the-art capabilities while meeting the highest standards for security and reliability.